Micron Document
Gemini Proxy


kennedy.gemi.dev kennedy.gemi.dev/docs/crawling.gmi
🔭 Notes on Crawling and Indexing


Kennedy creates its search index by crawling content within Geminispace. It does not crawl or index content from other protocols, such as Gopher or HTTP.

Crawler details

Kennedy crawls Geminispace using the following IP addresses:
* IPv4: 64.149.155.184
* IPv6: 2600:1700:1731:d0f:35a7:42d4:c71f:a02b

Crawler speed

Kennedy throttles itself and waits 1.5 seconds between making requests to the same IP address. This increases the amount of time it takes to crawl multiple capsules hosted from the same IP address, such as Flounder.online.

Robots.txt Support

Kennedy will respect sites that are using the simplified robots.txt protocol defined for Gemini.


Specifically, Kennedy will follow the Deny rules defined for the following user-agents:
* *
* indexer

Note: There are a number of robots.txt files in Geminispace which use rules outside of the simplified standard above. These include:
* Allow Rules
* Deny Rules with wildcard characters in the middle
* Crawl-Delay directives

Kennedy does not currently respect these rules.

Crawler Limits

Kennedy has the following limits:
* Will not download responses larger than 10 MB.
* Closes a connection if a URL takes more than 45 seconds to fully respond.

you're on kennedy.gemi.dev/docs/crawling.gmi